scientific image
Anyone can fake a scientific image with AI, tricking even academic journals – and undermining trust in science
A photograph of Earth glowing in deep space, the Moon's cratered horizon stretching across its foreground, caught many people's eyes in April 2026. Astronauts captured the image while aboard NASA's Artemis II mission, and like the famous Apollo 8 "Earthrise" image, the picture felt instantly real and inspiring for many. But when almost anyone can fabricate a visually similar image in seconds from a text prompt using artificial intelligence, how do people decide which image is real? The proliferation of AI-generated science images in public spaces is not simply a misinformation problem. As a researcher who studies visual science communication and public trust, I believe it also contributes to a crisis of trust in science in the age of AI, and the tools scientists have long relied on to establish visual credibility are losing their grip.
ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?
Zhang, Leixin, Eger, Steffen, Cheng, Yinjie, Zhai, Weihe, Belouadi, Jonas, Leiter, Christoph, Ponzetto, Simone Paolo, Moafian, Fahimeh, Zhao, Zhixue
Multimodal large language models (LLMs) have demonstrated impressive capabilities in generating high-quality images from textual instructions. However, their performance in generating scientific images--a critical application for accelerating scientific progress--remains underexplored. In this work, we address this gap by introducing ScImage, a benchmark designed to evaluate the multimodal capabilities of LLMs in generating scientific images from textual descriptions. ScImage assesses three key dimensions of understanding: spatial, numeric, and attribute comprehension, as well as their combinations, focusing on the relationships between scientific objects (e.g., squares, circles). We evaluate five models, GPT-4o, Llama, AutomaTikZ, Dall-E, and StableDiffusion, using two modes of output generation: code-based outputs (Python, TikZ) and direct raster image generation. Additionally, we examine four different input languages: English, German, Farsi, and Chinese. Our evaluation, conducted with 11 scientists across three criteria (correctness, relevance, and scientific accuracy), reveals that while GPT-4o produces outputs of decent quality for simpler prompts involving individual dimensions such as spatial, numeric, or attribute understanding in isolation, all models face challenges in this task, especially for more complex prompts.
Localization of Synthetic Manipulations in Western Blot Images
Manjunath, Anmol, Negroni, Viola, Mandelli, Sara, Moreira, Daniel, Bestagini, Paolo
Recent breakthroughs in deep learning and generative systems have significantly fostered the creation of synthetic media, as well as the local alteration of real content via the insertion of highly realistic synthetic manipulations. Local image manipulation, in particular, poses serious challenges to the integrity of digital content and societal trust. This problem is not only confined to multimedia data, but also extends to biological images included in scientific publications, like images depicting Western blots. In this work, we address the task of localizing synthetic manipulations in Western blot images. To discriminate between pristine and synthetic pixels of an analyzed image, we propose a synthetic detector that operates on small patches extracted from the image. We aggregate patch contributions to estimate a tampering heatmap, highlighting synthetic pixels out of pristine ones. Our methodology proves effective when tested over two manipulated Western blot image datasets, one altered automatically and the other manually by exploiting advanced AI-based image manipulation tools that are unknown at our training stage. We also explore the robustness of our method over an external dataset of other scientific images depicting different semantics, manipulated through unseen generation techniques.